Skip to content

[Model][NVIDIA] Route DSA models to the CUDA non-compiled path - #52861

Merged
WoosukKwon merged 16 commits into
vllm-project:mainfrom
WoosukKwon:codex/pr-49790-publish
Aug 19, 2026
Merged

WoosukKwon merged 16 commits into
vllm-project:mainfrom
WoosukKwon:codex/pr-49790-publish

Conversation

@WoosukKwon

@WoosukKwon WoosukKwon commented Aug 19, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Route DeepseekV32ForCausalLM, GlmMoeDsaForCausalLM, and their MTP draft
model through the CUDA vllm.models.deepseek_v32 implementation on every
NVIDIA GPU, while retaining the existing defaults elsewhere.

The optimized NVIDIA classes are currently unreachable from the model registry.
That leaves DeepSeek V3.2 and GLM-5.2 on the generic runner/graph path and
prevents their MTP draft model from using the matching implementation.

This PR:

  • dispatches DeepSeek V3.2 and GLM-5.2 DSA to the CUDA implementation on every
    NVIDIA GPU;
  • exports canonical DeepseekV32ForCausalLM and DeepseekV32MTP class names
    through the package platform entry point;
  • preserves DeepSeek V3.2's generic compiled MRV1 path on ROCm while keeping
    the explicit AMD modules available for opt-in use;
  • keeps individual SM100/SM103 optimizations capability-gated with existing
    fallbacks on earlier NVIDIA GPUs;
  • registers and routes DeepseekV32MTPModel with the main model;
  • defaults both DSA architectures to MRV2 with breakable full + piecewise CUDA
    graphs on NVIDIA, while retaining the ROCm compiled default;
  • supports unquantized BF16 and both standard and packed FP8 KV-cache forms;
  • removes the obsolete DeepSeek V3.2 MTP eager override;
  • selects CompilationMode.NONE for that graph path and preserves the explicit
    VLLM_USE_BREAKABLE_CUDAGRAPH=0 opt-out; and
  • adds registry and configuration coverage.

This supersedes #49790. That PR was automatically closed after its contributor
branch was synchronized to main; it now has no commits or changed files and
cannot be recovered by a maintainer push. The prepared changes were rebased onto
current main and published here instead.

Duplicate-work check

No open PR covers this NVIDIA CUDA default-routing scope. In particular, #51915 is
an opt-in ROCm/MXFP4 correctness path and explicitly leaves the default registry
route unchanged. The other open GLM-5.2/DSA results are backend- or
kernel-specific. This remains the NVIDIA DSA routing item tracked by #48597.

Test Plan

Run the affected configuration tests:

env -u VLLM_USE_BREAKABLE_CUDAGRAPH -u VLLM_USE_V2_MODEL_RUNNER \
  .venv/bin/python -m pytest tests/test_config.py -q

Check that the routed target and draft model classes import through the registry:

.venv/bin/python -m pytest \
  tests/models/test_registry.py::test_registry_imports \
  -k 'DeepseekV32ForCausalLM or GlmMoeDsaForCausalLM or DeepseekV32MTPModel' -q

Run pre-commit on all sixteen changed files.

Evaluate the full GSM8K set with TP=4 and MTP=3, then benchmark batch-size-1
serving on 4x NVIDIA GB200 with and without MTP. Both benchmark arms use the
same server command; the MTP arm adds the final --speculative-config option:

CUDA_VISIBLE_DEVICES=0,1,2,3 .venv/bin/vllm serve \
  nvidia/GLM-5.2-NVFP4 \
  --served-model-name GLM-5.2 \
  --revision aec724e8c7b8ee9db3b48c01c320f63f9cdaf8aa \
  --tensor-parallel-size 4 \
  --port 8300 \
  --kv-cache-dtype fp8_e4m3 \
  --max-model-len 16384 \
  --max-num-seqs 256 \
  --max-num-batched-tokens 16384 \
  --no-enable-prefix-caching \
  --gpu-memory-utilization 0.85 \
  --safetensors-load-strategy prefetch \
  --disable-uvicorn-access-log \
  --kernel-config \
    '{"ir_op_priority":{"rms_norm":["vllm_c","native"],"fused_add_rms_norm":["vllm_c","native"]},"enable_flashinfer_autotune":false}' \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

The serving environment has matching FlashInfer 0.6.17 packages:
flashinfer-python==0.6.17, flashinfer-cubin==0.6.17, and
flashinfer-jit-cache==0.6.17+cu130.

Test Result

tests/test_config.py:
199 passed, 14 warnings in 106.96s

focused registry imports:
3 passed, 375 deselected, 14 warnings in 7.04s

SM100 fused DSA norm/RoPE cache test (`auto`, `bfloat16`, `fp8`):
30 passed, 14 warnings in 13.19s

pre-commit run --files <all sixteen changed files>
All hooks passed, including ruff, formatting, mypy, SPDX, and config checks.

The import/configuration checks cover CUDA-wide routing without requiring a
specific compute capability and confirm that the non-CUDA package entry point resolves
to the generic target and MTP classes. The full model evaluation and performance
runs below use GLM-5.2 on GB200; DeepSeek V3.2 and pre-SM100 GPU runtime were not
benchmarked here.

Full GSM8K, 1,319 questions, 5-shot, temperature 0, seed 42, max output 256,
concurrency 100:

Exact-match accuracy:   94.768% (1,250/1,319)
Invalid responses:      0
Request errors:         0
256-token cap hits:     3
Questions/s:            29.818
Output tokens/s:        2,937.604
MTP mean accept length: 3.373
Draft-token acceptance: 79.104%

Batch-size-1 SPEED-Bench throughput_16k/low_entropy, 8,192 input tokens,
1,024 output tokens, concurrency 1, one warmup and three measured requests:

Metric No MTP MTP=3
Successful requests 3 3
Output tokens 3,072 3,072
Output throughput (includes TTFT) 130.17 tok/s 304.15 tok/s
Decode throughput (1 / TPOT) 135.09 tok/s 338.05 tok/s
Mean TPOT 7.40 ms 2.96 ms
Mean TTFT 293.90 ms 340.44 ms
Mean end-to-end latency 7,866.43 ms 3,366.65 ms
Draft-token acceptance - 78.90%
Mean accept length - 3.37

MTP=3 improves wall-clock output throughput by 2.337x (+133.7%) and
decode-only throughput by 2.502x (+150.2%). Startup logs confirmed MRV2,
automatic breakable FULL + PIECEWISE CUDA graphs, CompilationMode.NONE, sparse
FlashInfer MLA, FLASHINFER_TRTLLM NVFP4 MoE, allreduce_rms fusion, and
DeepseekV32MTPModel for the MTP arm.

A forced FlashMLA smoke and batch-size-1 benchmark also passed on 4x GB200 with the real NVFP4 checkpoint. The server used --attention-backend FLASHMLA_SPARSE --kv-cache-dtype fp8_ds_mla, TP=4, V2 model runner, and breakable full + piecewise CUDA graphs. Logs confirmed FLASHMLA_SPARSE attention and FLASHINFER_TRTLLM NVFP4 MoE. A 2,048-input/1,024-output SPEED-Bench request achieved 114.01 output tok/s including TTFT, 8.75 ms TPOT (about 114.3 decode tok/s), and 34.48 ms TTFT. The request completed successfully with no runtime JIT warning.

Compile-test cleanup:

.venv/bin/python -m pytest tests/compile/h100/test_startup.py --collect-only -q
6 tests collected; no DeepSeek V3.2 case

.venv/bin/python -m pytest tests/compile/fusions_e2e/test_tp1_quant.py \
  tests/compile/fusions_e2e/test_tp2_ar_rms.py --collect-only -q
568 tests collected; no DeepSeek V3.2 NVFP4 case

.venv/bin/python -m py_compile <all five compile-test files changed>
Passed

pre-commit run --files <all five compile-test files changed>
All applicable hooks passed.

DeepSeek V3.2 was removed from the H100 compile-startup suite and from the CUDA NVFP4 fusion-E2E suite because the NVIDIA path defaults to breakable CUDA graphs with compilation disabled. The generic ROCm/AITER compiler-pass tests remain because AMD intentionally stays on the compiled path.

AI assistance was used to prepare the rebase/default-path changes, debug the
benchmark environment, run tests and evaluations, and draft this description.
The human submitter must review every changed line and the evidence above and
own the change end-to-end before marking this draft ready for review.


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR is described and linked to the existing tracker.
  • Test commands are provided.
  • Test and model-evaluation results are provided.
  • No documentation update is required; this changes internal routing and defaults.

BEFORE SUBMITTING, PLEASE READ https://docs.vllm.ai/en/latest/contributing

zhou9402 and others added 4 commits August 19, 2026 01:30
Route GlmMoeDsaForCausalLM and the matching MTP architecture to the
optimized deepseek_v32 implementation on SM100-family devices, keeping the
generic deepseek_v2 fallback everywhere else. Default the KV cache to FP8
since this implementation requires a sparse FP8 cache.

Split out of vllm-project#48597 (reverted by vllm-project#49768).

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Peiyuan Zhou <peiyuanzhou1994@gmail.com>
mypy rejects rebinding a class name with a plain assignment ("Cannot assign to
a type"), which the other branches of this dispatch bind by import. Use an
import alias so every branch defines the name the same way.

Co-authored-by: Claude Opus 5
Signed-off-by: Peiyuan Zhou <peiyuanzhou1994@gmail.com>
Routing GlmMoeDsaForCausalLM here changed what a default launch does. This
implementation asserted an fp8 KV cache, so an unset --kv-cache-dtype was
rewritten to fp8 from inside a per-layer constructor: the engine went on
reporting kv_cache_dtype=auto while allocating an fp8 cache (2,935,232 KV
tokens against main's 1,509,888 for the same launch), and the "Using ... data
type to store kv cache" line never printed.

The assert was stricter than anything below it needs. fused_norm_rope already
writes an unquantized cache when the dtype is not fp8, fused_q already emits
the bf16 (ql_nope, q_pe) query that the fp8_ds_mla layout uses, FlashInfer
sparse accepts that tuple, and the ROCm subclass already derives the same two
flags from the cache dtype. Only the query form (taken from the backend
capability rather than the cache) and the unconditional fp8 view of the paged
cache assumed fp8; both now come from the dtype, and nothing rewrites
cache_dtype.

Measured on 8xB300, GLM-5.2 block-fp8, TP8, MTP=5, 8192-in/1024-out at
concurrency 1. Default launch: 1,513,024 KV tokens (bf16, matching main's
1,509,888), 362 tok/s vs main's 340 for the same config, GSM8K 0.950.
--kv-cache-dtype fp8_e4m3 is unchanged at 2,935,232 tokens and 403 tok/s,
GSM8K 0.938.

Co-authored-by: Claude Opus 5
Signed-off-by: Peiyuan Zhou <peiyuanzhou1994@gmail.com>
Choose Model Runner V2 for GLM-5.2 and auto-enable the non-compiled breakable CUDA graph path for the model and MTP architectures on every platform. The optimized SM100 model routing remains hardware-specific, while this serving default does not.

Cover both NVFP4 and FP8 checkpoints with and without MTP.

Co-authored-by: OpenAI Codex <codex@openai.com>

Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
@mergify mergify Bot added deepseek Related to DeepSeek models new-model Requests to new models nvidia speculative-decoding labels Aug 19, 2026
GLM-5.2 now defaults to the non-compiled MRV2 path, so route every CUDA device through the NVIDIA deepseek_v32 implementation. Capability-specific kernels continue to gate themselves and fall back when unavailable.

Drop the standalone routing and KV-cache-form test files.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
@WoosukKwon WoosukKwon changed the title [Model][NVIDIA] Route GLM-5.2 DSA to the SM100 implementation [Model][NVIDIA] Route GLM-5.2 DSA to the CUDA non-compiled path Aug 19, 2026
WoosukKwon and others added 2 commits August 19, 2026 03:41
Use the deepseek_v32 main and MTP implementations for DeepSeek V3.2, including NVIDIA NVFP4 checkpoints. Default both architectures to MRV2 with breakable CUDA graphs and remove the obsolete MTP eager override.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Resolve the registry through CUDA-specific aliases so NVIDIA uses the deepseek_v32 implementation while ROCm, XPU, and CPU retain their existing defaults. Keep the explicit AMD package exports available for opt-in use.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
@WoosukKwon WoosukKwon changed the title [Model][NVIDIA] Route GLM-5.2 DSA to the CUDA non-compiled path [Model][NVIDIA] Route DSA models to the CUDA non-compiled path Aug 19, 2026
WoosukKwon and others added 5 commits August 19, 2026 03:53
Move the CUDA-or-generic registry dispatch into a separate module that exports DeepseekV32ForCausalLM and DeepseekV32MTP without prefixed aliases. Keep ROCm on the generic compiled MRV1 path without automatic breakable graphs while preserving the opt-in AMD package exports.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
@WoosukKwon

Copy link
Copy Markdown
Collaborator Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #84554 for commit e73990b0a406.

@WoosukKwon
WoosukKwon merged commit b09bd69 into vllm-project:main Aug 19, 2026
124 checks passed
@github-project-automation github-project-automation Bot moved this to Done in NVIDIA Aug 19, 2026
@WoosukKwon
WoosukKwon deleted the codex/pr-49790-publish branch August 19, 2026 16:12
@WoosukKwon

Copy link
Copy Markdown
Collaborator Author

Results

NVFP4 TP4, dummy weights, no MTP

This is pure tensor parallelism (tensor_parallel_size=4, DP=1, EP disabled).
Both variants completed every request with the same 0.85 memory target and all
83 default CUDA graph capture sizes.

Concurrency Baseline output tok/s Candidate output tok/s Delta
1 111.63 129.96 +16.42%
8 688.56 791.61 +14.97%
32 1,565.80 1,738.79 +11.05%
256 2,244.52 2,430.66 +8.29%
1024 2,250.54 2,444.41 +8.61%

Candidate mean TPOT was also lower at every point: 7.41/8.76/15.51/43.34/
45.32 ms versus baseline 8.65/10.19/17.30/46.57/48.71 ms. This matrix shows
no performance regression and passes the required TP4 concurrency-1 threshold.

NVFP4 TP4, real weights, MTP3

This is also pure tensor parallelism (DP=1, EP disabled). Both variants loaded
the real 47-shard checkpoint, used the sparse MTP architecture, and completed
all requests without failures.

Concurrency Baseline output tok/s Candidate output tok/s Delta
1 262.84 309.54 +17.77%
8 979.18 1,057.14 +7.96%
32 1,781.36 1,891.51 +6.18%
256 2,294.24 2,485.39 +8.33%
1024 2,312.33 2,496.11 +7.95%

Candidate mean TPOT was lower at every point: 2.94/6.57/14.32/37.60/39.27 ms
versus baseline 3.48/7.15/15.17/40.38/42.05 ms. Acceptance was about 79--80%
for the saturated points. Greedy output text is not byte-identical between the
two implementations; the paired real-weight non-speculative evaluation below
provides the output-quality check.

NVFP4 DEP4, dummy weights, no MTP

These runs use TP1 x DP4 with expert parallelism enabled. They are therefore
not subject to the pure-TP4 concurrency-1 throughput gate. Both variants
completed all requests without failures.

Concurrency Baseline output tok/s Candidate output tok/s Delta
1 61.63 68.25 +10.74%
8 441.88 451.51 +2.18%
32 1,363.61 1,594.39 +16.92%
256 2,852.79 3,589.51 +25.82%
1024 2,988.89 3,667.67 +22.71%

Candidate had more KV capacity per replica (367,616 versus 284,544 tokens)
because it did not retain compiled graph artifacts. At concurrency 256/1024,
that allowed more simultaneous work and substantially improved throughput and
TTFT, while mean TPOT was slightly higher (40.16/41.93 ms candidate versus
39.41/39.90 ms baseline). This is a saturation scheduling tradeoff rather than
an end-to-end regression; lower concurrencies improved both throughput and
TPOT. The DP stats coordinator also emitted out-of-order-step warnings on both
variants without request failures.

NVFP4 DEP4, real weights, MTP3

These runs use TP1 x DP4 with expert parallelism enabled and are not subject
to the pure-TP4 concurrency-1 threshold. Both variants use a matched 1.5 GiB
KV-cache reservation per GPU, the automatic NVFP4 target MoE backend, and the
draft-only Triton override described above. All completed points have zero
failed requests.

Concurrency Baseline output tok/s Candidate output tok/s Delta
1 160.03 165.86 +3.64%
8 685.85 745.49 +8.70%
32 870.99 880.50 +1.09%
256 864.60 922.61 +6.71%
1024 863.07 928.86 +7.62%

Candidate mean TPOT is lower at concurrency 256/1024: 12.65/12.59 ms versus
13.41/13.51 ms for baseline. MTP acceptance is closely matched: 79.53/79.69%
candidate versus 79.39/80.02% baseline. Every point completed with zero failed
requests.

FP8 TP8, dummy weights, no MTP

This is pure tensor parallelism (tensor_parallel_size=8, DP=1, EP disabled).
Both variants completed every request with the same 0.85 memory target and all
83 default CUDA graph capture sizes.

Concurrency Baseline output tok/s Candidate output tok/s Delta
1 89.47 97.60 +9.09%
8 503.21 545.85 +8.47%
32 1,075.94 1,152.79 +7.14%
256 1,887.46 2,040.56 +8.11%
1024 1,957.64 2,110.54 +7.81%

The candidate is faster at every measured concurrency and every point has zero
failed requests. The absence of an underperforming point means there is no FP8
performance gap to attribute with a profiler trace.

FP8 TP8, real weights, MTP3

This is pure tensor parallelism (DP=1, EP disabled). Both variants loaded the
real 141-shard checkpoint, used MTP3, retained all 83 default capture sizes,
and completed every request without failures.

Concurrency Baseline output tok/s Candidate output tok/s Delta
1 228.96 250.68 +9.49%
8 790.04 821.65 +4.00%
32 1,547.02 1,572.03 +1.62%
256 2,203.24 2,342.57 +6.32%
1024 2,266.40 2,458.07 +8.46%

Candidate mean TPOT is lower at every point: 3.67/8.50/17.32/54.39/55.59 ms
versus baseline 4.02/9.00/17.76/56.92/59.21 ms. MTP acceptance is closely
matched at about 79--81%. The candidate is faster at every concurrency, so no
profiler trace is warranted for this matrix.

FP8 DEP8, dummy weights, no MTP

These runs use TP1 x DP8 with expert parallelism enabled, four local DP ranks
per node, and the Ray internal load balancer. They are not subject to the
pure-TP concurrency-1 throughput gate. Both variants completed every request
with all 83 default capture sizes.

Concurrency Baseline output tok/s Candidate output tok/s Delta
1 42.95 45.48 +5.89%
8 302.31 320.71 +6.09%
32 998.29 1,048.39 +5.02%
256 3,106.51 3,292.76 +6.00%
1024 3,077.41 3,522.77 +14.47%

Candidate KV capacity is 461,440 tokens per replica versus 336,064 for the
compiled baseline. The candidate is faster at every concurrency and all points
have zero failed requests. Its higher capacity increases saturated throughput
at concurrency 1024 while also increasing mean TPOT there (103.02 ms versus
85.99 ms), a queue/residency tradeoff rather than an end-to-end regression.

FP8 DEP8, real weights, MTP3

These runs use TP1 x DP8 with expert parallelism enabled, four local DP ranks
per node, and the Ray internal load balancer. They use the real checkpoint,
MTP3, a matched 1.5 GiB KV allocation (33,216 tokens per replica), and all 83
default capture sizes. The target MoE uses DeepGEMM; the draft-only override
uses Triton. This DEP topology is not subject to the pure-TP4 throughput gate.

Concurrency Baseline output tok/s Candidate output tok/s Delta
1 111.79 124.41 +11.29%
8 578.66 627.71 +8.48%
32 954.78 1,026.77 +7.54%
256 1,015.43 1,083.35 +6.69%
1024 1,017.86 1,083.16 +6.42%

Both variants completed every request. Candidate mean TPOT is lower at every
point: 7.47/11.36/20.49/21.62/21.83 ms versus baseline
8.31/12.14/21.61/23.15/23.12 ms. MTP acceptance is closely matched at about
79--83%. The candidate is faster at every concurrency, so no profiler trace is
warranted for this matrix.

Real-weight output-quality check

A paired non-speculative TP4 evaluation used the real NVFP4 checkpoint, MRV2,
temperature 0, seed 42, 100 fixed GSM8K questions with five shots, and a
256-token output limit. EP and MTP were disabled. Both servers retained
FULL_AND_PIECEWISE and the 83 default capture sizes.

Variant GSM8K accuracy Invalid responses Output tokens
Baseline 95/100 0/100 9,601
Candidate 96/100 1/100 9,625

The candidate has no accuracy regression on this fixed sample. Its single
invalid response is offset by one additional correct answer and does not
indicate a systematic malformed-output problem.

zufangzhu pushed a commit to zufangzhu/vllm that referenced this pull request Aug 24, 2026
weijinqian0 pushed a commit to vllm-project/vllm-ascend that referenced this pull request Aug 24, 2026
### What this PR does / why we need it?
The PR adapts vllm-ascend for compatibility with the latest vLLM main
(commit `ba07e4a4`).

| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
| `.github/vllm-main-verified.commit` | — | Updated verified main commit
hash from `cdc4824a21` to `ba07e4a48` |
| `tests/e2e/conftest.py` |
[vllm#53272](vllm-project/vllm#53272) — upstream
plans to remove native Hunyuan V1/VL;
[vllm#51665](vllm-project/vllm#51665) — dropped
HunYuanVL `lm_head` workaround | Added `skip` condition for HunyuanVL
e2e when `vllm_version_is("0.27.1")` is False (vLLM main) |
| `tests/ut/core/test_profiling_chunk.py` | — vLLM main
`Scheduler.__init__` reads `model_config.uses_mrope`, which infinitely
recurses on a bare MagicMock | Version-gated: sets
`type(model_config).uses_mrope = PropertyMock(return_value=False)` on
main |
| `tests/ut/core/test_recompute_scheduler.py` | — vLLM main
`Scheduler.add_request` reads `spec_decode_metrics_level` |
Version-gated: sets `scheduler.spec_decode_metrics_level = "none"` on
main |
| `tests/ut/patch/platform/test_patch_structured_output.py` | — Upstream
changed structured output validation error type from `ValueError` to
`VLLMValidationError` | Version-gated: `error_type = ValueError if
vllm_version_is("0.27.1") else VLLMValidationError` used in all three
fake validation functions and `pytest.raises` |
| `vllm_ascend/attention/attention_v1.py` |
[vllm#52839](vllm-project/vllm#52839) — moved
`pcp.py` from `vllm.model_executor.layers.attention.pcp` to
`vllm.v1.attention.ops.pcp` | Version-gated
`_gather_prefill_cache_inputs` import path |
| `vllm_ascend/compilation/acl_graph.py` |
[vllm#49134](vllm-project/vllm#49134) —
`get_current_vllm_config()` now raises `AssertionError` when called
outside `set_current_vllm_config()` context | `update_full_graph_params`
version-gated: main branch wraps `get_impl_cls()`/`update_graph_params`
in `with set_current_vllm_config(vllm_config):` |
| `vllm_ascend/models/deepseek_mtp.py` |
[vllm#53106](vllm-project/vllm#53106) — removed
`skip_prefixes` kwarg from `AutoWeightsLoader.__init__` |
`AscendGlmMoeDsaForCausalLM.load_weights` version-gated: 0.27.1 uses
`skip_prefixes=["rot."]`; main uses
`WeightsMapper(orig_to_new_prefix={"rot.": None})` passed via
`load_weights(weights, mapper=mapper)` |
| `vllm_ascend/spec_decode/llm_base_proposer.py` |
[vllm#52861](vllm-project/vllm#52861) — added
`DeepseekV32MTPModel` to MTP architecture set in `model_returns_tuple()`
| `model_returns_tuple` version-gated: 0.27.1 checks
`{"DeepSeekMTPModel", "KimiK3MTPModel"}`; main also includes
`"DeepseekV32MTPModel"` |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` |
[vllm#52188](vllm-project/vllm#52188) — added
`cp_rank`, `CP_SIZE`, `CP_INTERLEAVE` params to
`_prepare_dflash_inputs_kernel` for DCP support | Entire
`_prepare_dflash_inputs_kernel_ascend` kernel duplicated under
`vllm_version_is("0.27.1")` gate: 0.27.1 uses 30 pos + 3 constexpr (no
DCP params); main uses 31 pos + 5 constexpr |
### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@cdc4824

---------

Signed-off-by: hfadzxy <starmoon_zhang@163.com>
yiminghub2024 pushed a commit to yiminghub2024/vllm-ascend that referenced this pull request Aug 25, 2026
### What this PR does / why we need it?
The PR adapts vllm-ascend for compatibility with the latest vLLM main
(commit `ba07e4a4`).

| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
| `.github/vllm-main-verified.commit` | — | Updated verified main commit
hash from `cdc4824a21` to `ba07e4a48` |
| `tests/e2e/conftest.py` |
[vllm#53272](vllm-project/vllm#53272) — upstream
plans to remove native Hunyuan V1/VL;
[vllm#51665](vllm-project/vllm#51665) — dropped
HunYuanVL `lm_head` workaround | Added `skip` condition for HunyuanVL
e2e when `vllm_version_is("0.27.1")` is False (vLLM main) |
| `tests/ut/core/test_profiling_chunk.py` | — vLLM main
`Scheduler.__init__` reads `model_config.uses_mrope`, which infinitely
recurses on a bare MagicMock | Version-gated: sets
`type(model_config).uses_mrope = PropertyMock(return_value=False)` on
main |
| `tests/ut/core/test_recompute_scheduler.py` | — vLLM main
`Scheduler.add_request` reads `spec_decode_metrics_level` |
Version-gated: sets `scheduler.spec_decode_metrics_level = "none"` on
main |
| `tests/ut/patch/platform/test_patch_structured_output.py` | — Upstream
changed structured output validation error type from `ValueError` to
`VLLMValidationError` | Version-gated: `error_type = ValueError if
vllm_version_is("0.27.1") else VLLMValidationError` used in all three
fake validation functions and `pytest.raises` |
| `vllm_ascend/attention/attention_v1.py` |
[vllm#52839](vllm-project/vllm#52839) — moved
`pcp.py` from `vllm.model_executor.layers.attention.pcp` to
`vllm.v1.attention.ops.pcp` | Version-gated
`_gather_prefill_cache_inputs` import path |
| `vllm_ascend/compilation/acl_graph.py` |
[vllm#49134](vllm-project/vllm#49134) —
`get_current_vllm_config()` now raises `AssertionError` when called
outside `set_current_vllm_config()` context | `update_full_graph_params`
version-gated: main branch wraps `get_impl_cls()`/`update_graph_params`
in `with set_current_vllm_config(vllm_config):` |
| `vllm_ascend/models/deepseek_mtp.py` |
[vllm#53106](vllm-project/vllm#53106) — removed
`skip_prefixes` kwarg from `AutoWeightsLoader.__init__` |
`AscendGlmMoeDsaForCausalLM.load_weights` version-gated: 0.27.1 uses
`skip_prefixes=["rot."]`; main uses
`WeightsMapper(orig_to_new_prefix={"rot.": None})` passed via
`load_weights(weights, mapper=mapper)` |
| `vllm_ascend/spec_decode/llm_base_proposer.py` |
[vllm#52861](vllm-project/vllm#52861) — added
`DeepseekV32MTPModel` to MTP architecture set in `model_returns_tuple()`
| `model_returns_tuple` version-gated: 0.27.1 checks
`{"DeepSeekMTPModel", "KimiK3MTPModel"}`; main also includes
`"DeepseekV32MTPModel"` |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` |
[vllm#52188](vllm-project/vllm#52188) — added
`cp_rank`, `CP_SIZE`, `CP_INTERLEAVE` params to
`_prepare_dflash_inputs_kernel` for DCP support | Entire
`_prepare_dflash_inputs_kernel_ascend` kernel duplicated under
`vllm_version_is("0.27.1")` gate: 0.27.1 uses 30 pos + 3 constexpr (no
DCP params); main uses 31 pos + 5 constexpr |
### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@cdc4824

---------

Signed-off-by: hfadzxy <starmoon_zhang@163.com>
frankie-ys pushed a commit to Csrayz/vllm-ascend that referenced this pull request Aug 26, 2026
### What this PR does / why we need it?
The PR adapts vllm-ascend for compatibility with the latest vLLM main
(commit `ba07e4a4`).

| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
| `.github/vllm-main-verified.commit` | — | Updated verified main commit
hash from `cdc4824a21` to `ba07e4a48` |
| `tests/e2e/conftest.py` |
[vllm#53272](vllm-project/vllm#53272) — upstream
plans to remove native Hunyuan V1/VL;
[vllm#51665](vllm-project/vllm#51665) — dropped
HunYuanVL `lm_head` workaround | Added `skip` condition for HunyuanVL
e2e when `vllm_version_is("0.27.1")` is False (vLLM main) |
| `tests/ut/core/test_profiling_chunk.py` | — vLLM main
`Scheduler.__init__` reads `model_config.uses_mrope`, which infinitely
recurses on a bare MagicMock | Version-gated: sets
`type(model_config).uses_mrope = PropertyMock(return_value=False)` on
main |
| `tests/ut/core/test_recompute_scheduler.py` | — vLLM main
`Scheduler.add_request` reads `spec_decode_metrics_level` |
Version-gated: sets `scheduler.spec_decode_metrics_level = "none"` on
main |
| `tests/ut/patch/platform/test_patch_structured_output.py` | — Upstream
changed structured output validation error type from `ValueError` to
`VLLMValidationError` | Version-gated: `error_type = ValueError if
vllm_version_is("0.27.1") else VLLMValidationError` used in all three
fake validation functions and `pytest.raises` |
| `vllm_ascend/attention/attention_v1.py` |
[vllm#52839](vllm-project/vllm#52839) — moved
`pcp.py` from `vllm.model_executor.layers.attention.pcp` to
`vllm.v1.attention.ops.pcp` | Version-gated
`_gather_prefill_cache_inputs` import path |
| `vllm_ascend/compilation/acl_graph.py` |
[vllm#49134](vllm-project/vllm#49134) —
`get_current_vllm_config()` now raises `AssertionError` when called
outside `set_current_vllm_config()` context | `update_full_graph_params`
version-gated: main branch wraps `get_impl_cls()`/`update_graph_params`
in `with set_current_vllm_config(vllm_config):` |
| `vllm_ascend/models/deepseek_mtp.py` |
[vllm#53106](vllm-project/vllm#53106) — removed
`skip_prefixes` kwarg from `AutoWeightsLoader.__init__` |
`AscendGlmMoeDsaForCausalLM.load_weights` version-gated: 0.27.1 uses
`skip_prefixes=["rot."]`; main uses
`WeightsMapper(orig_to_new_prefix={"rot.": None})` passed via
`load_weights(weights, mapper=mapper)` |
| `vllm_ascend/spec_decode/llm_base_proposer.py` |
[vllm#52861](vllm-project/vllm#52861) — added
`DeepseekV32MTPModel` to MTP architecture set in `model_returns_tuple()`
| `model_returns_tuple` version-gated: 0.27.1 checks
`{"DeepSeekMTPModel", "KimiK3MTPModel"}`; main also includes
`"DeepseekV32MTPModel"` |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` |
[vllm#52188](vllm-project/vllm#52188) — added
`cp_rank`, `CP_SIZE`, `CP_INTERLEAVE` params to
`_prepare_dflash_inputs_kernel` for DCP support | Entire
`_prepare_dflash_inputs_kernel_ascend` kernel duplicated under
`vllm_version_is("0.27.1")` gate: 0.27.1 uses 30 pos + 3 constexpr (no
DCP params); main uses 31 pos + 5 constexpr |
### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@cdc4824

---------

Signed-off-by: hfadzxy <starmoon_zhang@163.com>
Rohan138 added a commit to ROCm/vllm that referenced this pull request Aug 28, 2026
… cudagraph

vLLM vllm-project#52861 routed the DSA architectures onto the non-compiled V2 model
runner / breakable-cudagraph path, but GlmMoeDsaForCausalLM was left out of
the ROCm carve-out, flipping GLM-5.2 to MRV2 + breakable cudagraphs and
regressing batch-1 decode TPOT by ~30-37% on gfx950 (FP8 and MXFP4).

Add GlmMoeDsaForCausalLM to ROCM_DEFAULT_MRV1_ARCHITECTURES so GLM-5.2 stays
on the compiled MRV1 path, and default breakable cudagraphs off entirely on
ROCm (they regress performance today); VLLM_USE_BREAKABLE_CUDAGRAPH=1 still
forces them on.

Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Lethobenthos20 pushed a commit to Lethobenthos20/vllm-ascend that referenced this pull request Sep 4, 2026
### What this PR does / why we need it?
The PR adapts vllm-ascend for compatibility with the latest vLLM main
(commit `ba07e4a4`).

| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
| `.github/vllm-main-verified.commit` | — | Updated verified main commit
hash from `cdc4824a21` to `ba07e4a48` |
| `tests/e2e/conftest.py` |
[vllm#53272](vllm-project/vllm#53272) — upstream
plans to remove native Hunyuan V1/VL;
[vllm#51665](vllm-project/vllm#51665) — dropped
HunYuanVL `lm_head` workaround | Added `skip` condition for HunyuanVL
e2e when `vllm_version_is("0.27.1")` is False (vLLM main) |
| `tests/ut/core/test_profiling_chunk.py` | — vLLM main
`Scheduler.__init__` reads `model_config.uses_mrope`, which infinitely
recurses on a bare MagicMock | Version-gated: sets
`type(model_config).uses_mrope = PropertyMock(return_value=False)` on
main |
| `tests/ut/core/test_recompute_scheduler.py` | — vLLM main
`Scheduler.add_request` reads `spec_decode_metrics_level` |
Version-gated: sets `scheduler.spec_decode_metrics_level = "none"` on
main |
| `tests/ut/patch/platform/test_patch_structured_output.py` | — Upstream
changed structured output validation error type from `ValueError` to
`VLLMValidationError` | Version-gated: `error_type = ValueError if
vllm_version_is("0.27.1") else VLLMValidationError` used in all three
fake validation functions and `pytest.raises` |
| `vllm_ascend/attention/attention_v1.py` |
[vllm#52839](vllm-project/vllm#52839) — moved
`pcp.py` from `vllm.model_executor.layers.attention.pcp` to
`vllm.v1.attention.ops.pcp` | Version-gated
`_gather_prefill_cache_inputs` import path |
| `vllm_ascend/compilation/acl_graph.py` |
[vllm#49134](vllm-project/vllm#49134) —
`get_current_vllm_config()` now raises `AssertionError` when called
outside `set_current_vllm_config()` context | `update_full_graph_params`
version-gated: main branch wraps `get_impl_cls()`/`update_graph_params`
in `with set_current_vllm_config(vllm_config):` |
| `vllm_ascend/models/deepseek_mtp.py` |
[vllm#53106](vllm-project/vllm#53106) — removed
`skip_prefixes` kwarg from `AutoWeightsLoader.__init__` |
`AscendGlmMoeDsaForCausalLM.load_weights` version-gated: 0.27.1 uses
`skip_prefixes=["rot."]`; main uses
`WeightsMapper(orig_to_new_prefix={"rot.": None})` passed via
`load_weights(weights, mapper=mapper)` |
| `vllm_ascend/spec_decode/llm_base_proposer.py` |
[vllm#52861](vllm-project/vllm#52861) — added
`DeepseekV32MTPModel` to MTP architecture set in `model_returns_tuple()`
| `model_returns_tuple` version-gated: 0.27.1 checks
`{"DeepSeekMTPModel", "KimiK3MTPModel"}`; main also includes
`"DeepseekV32MTPModel"` |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` |
[vllm#52188](vllm-project/vllm#52188) — added
`cp_rank`, `CP_SIZE`, `CP_INTERLEAVE` params to
`_prepare_dflash_inputs_kernel` for DCP support | Entire
`_prepare_dflash_inputs_kernel_ascend` kernel duplicated under
`vllm_version_is("0.27.1")` gate: 0.27.1 uses 30 pos + 3 constexpr (no
DCP params); main uses 31 pos + 5 constexpr |
### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.27.1
- vLLM main:
vllm-project/vllm@cdc4824

---------

Signed-off-by: hfadzxy <starmoon_zhang@163.com>
sychen52 added a commit to sychen52/vllm that referenced this pull request Sep 11, 2026
 vllm-project#52861 routed the DSA models to a fused norm+rope Triton kernel that
writes the MLA KV cache itself and only supports fp8_ds_mla. vllm-project#51724
added nvfp4_ds_mla after that and the rebase missed it.

Teach the fused kernel the nvfp4_ds_mla layout.

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deepseek Related to DeepSeek models new-model Requests to new models nvidia ready ONLY add when PR is ready to merge/full CI is needed speculative-decoding

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants